Papers with spatial and temporal representations
VideoQA-TA: Temporal-Aware Multi-Modal Video Question Answering (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods for video question answering align visual or textual features directly with large language models, limiting the deep semantic association between modalities and hindering a comprehensive understanding of interactions within spatial and temporal contexts. |
| Approach: | They propose a temporal-aware framework for multi-modal video question answering that aligns videos and questions at fine-grained levels. |
| Outcome: | The proposed framework improves reasoning ability and accuracy of videoQA by aligning videos and questions at fine-grained levels. |
Multi-Channel Spatio-Temporal Transformer for Sign Language Production (2024.lrec-main)
Copied to clipboard
| Challenge: | Sign language production models ignore structural correlations between channels and use multi-channel spatial attention to capture correlations across channels. |
| Approach: | They propose a novel approach to transform sign language into a unified feature representation using multi-channel spatial attention and temporal attention to learn sequential dependencies for each channel over time. |
| Outcome: | The proposed model outperforms state-of-the-art models on two sign language datasets from diverse cultures. |